Skip to main content

20. Normalization

Normalization keeps the values flowing through the network in a stable range.

It can make training faster and more stable.

Common normalization techniques include:

  • Batch Normalization
  • Layer Normalization

Batch Normalization​

Normalizes activations using statistics from the current batch.

self.batchnorm = nn.BatchNorm1d(10)

Useful in many feed-forward and CNN networks.

Layer Normalization​

Normalizes across the features of each individual sample.

self.layernorm = nn.LayerNorm(10)

Commonly used in Transformers.

Main Difference​

Batch NormalizationLayer Normalization
Normalizes across the batchNormalizes across features
Depends on batch statisticsWorks independently for each sample
Common in CNNsCommon in Transformers

Actual Implementation​

import torch
import torch.nn as nn

class NeuralNetwork(nn.Module):
def __init__(self):
super().__init__()

self.linear1 = nn.Linear(2, 10)
self.batchnorm = nn.BatchNorm1d(10)
self.linear2 = nn.Linear(10, 1)

def forwardpass(self, x):
x = self.linear1(x)
x = self.batchnorm(x)
x = torch.relu(x)
x = self.linear2(x)

return x

neuralnetwork = NeuralNetwork()

inputs = torch.tensor([
[18.0, 28.0],
[19.0, 29.0],
[20.0, 30.0],
[21.0, 31.0]
])

output = neuralnetwork.forwardpass(inputs)

print(output)

Quick Difference​

BatchNorm → normalize across the batch
LayerNorm → normalize across features of each sample